Papers by Verginica Barbu Mititelu
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)
Copied to clipboard
Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadić, Vanja Štefanec, Maciej Ogrodniczuk, Bartłomiej Nitoń, Piotr Pęzik, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan, Dan Tufiș, Radovan Garabík, Simon Krek, Andraž Repar
| Challenge: | The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains. |
| Approach: | They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains . |
| Outcome: | The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed . |
The Reference Corpus of the Contemporary Romanian Language (CoRoLa) (L18-1)
Copied to clipboard
| Challenge: | a four-year project focused on the creation of a big corpus for contemporary Romanian language is underway . the corpus is the largest publicly available corpus of contemporary Romania . |
| Approach: | a four-year project is focusing on the creation of a big corpus for Romanian language . the corpus is the largest publicly available corpus of the language based in the country . authors propose to use the corpora as a tool to query and listen to the results . |
| Outcome: | a four-year project has created the largest publicly available corpus of Romanian language . the corpus is the result of a project focused on the creation of 'corola.racai.ro' the written component contains 1,257,752,812 tokens, distributed in several languages . |
Ensemble Romanian Dependency Parsing with Neural Networks (L18-1)
Copied to clipboard
| Challenge: | SSPR is a Python 3.5 application based on the Microsoft Cognitive Toolkit 2.0 Python API. |
| Approach: | a Python 3.5 application is based on the Microsoft Cognitive Toolkit 2.0 Python API. |
| Outcome: | SSPR outperforms the best individual parser at the CONLL 2017 dependency parsing shared task. |
Aligning the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs (2022.lrec-1)
Copied to clipboard
| Challenge: | Among the language resources for Romanian, there are ones that describe the syntactic and semantic aspects of the language. |
| Approach: | They propose to align two language resources for Romanian: the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs. |
| Outcome: | The proposed alignments identify morpho-syntactic annotation mistakes, incomplete valence frames or missing ones. |
The MARCELL Legislative Corpus (2020.lrec-1)
Copied to clipboard
Tamás Váradi, Svetla Koeva, Martin Yamalov, Marko Tadić, Bálint Sass, Bartłomiej Nitoń, Maciej Ogrodniczuk, Piotr Pęzik, Verginica Barbu Mititelu, Radu Ion, Elena Irimia, Maria Mitrofan, Vasile Păiș, Dan Tufiș, Radovan Garabík, Simon Krek, Andraz Repar, Matjaž Rihtar, Janez Brank
| Challenge: | MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification. |
| Approach: | They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents . |
| Outcome: | The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents. |